Papers with reward objective
Countering Reward Over-Optimization in LLM with Demonstration-Guided Reinforcement Learning (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches address ROO by adding KL regularization, requiring computationally expensive hyperparameter tuning. |
| Approach: | They propose a reinforcement learning approach that leverages human demonstrations and a reward model to recalibrate the reward objective. |
| Outcome: | The proposed approach achieves comparable performance to carefully tuned baselines while mitigating ROO in three RL language tasks. |